feat(config): minimum-confidence policy and config version stamp - #178
Merged
Merged
Conversation
Three design specs and their task-by-task implementation plans, covering the accuracy and performance work queued after 1.9.0: - tree-sitter scan scale: the tree path is capped at 64 files / 256 KiB per scan because trees are retained for the whole scan, so on any real repository most files silently fall back to the regex path the spike measured at 60%/54.5% precision. Replaces retention with a bounded resident cache plus heap-aware parse admission, and counts every fallback. - cross-file taint: Go taint analysis resolves calls only within one file, so a flow through a helper in another file is a false negative. Adds a package-qualified summary index iterated to a fixed point in reverse topological package order, with a package-closure cache fingerprint so cross-file findings cannot go stale. - confidence policy: findings carry a confidence nothing consumes. Makes it a configurable minimum threshold with full accounting, optional demotion, and confidence-aware ordering. Each plan is red-first per task and carries its own acceptance criteria. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Findings have always carried a confidence of high, medium, or low, set by 177 call sites, but nothing consumed it: no filtering, no ordering, no effect on status. Structural analyses (source-to-sink taint, the tree-sitter rule paths) set high while regex line scans leave it unspecified, so a team could not ask for only the trustworthy half without disabling rules outright or baselining the noise. checks.min_confidence.default drops findings below a level, and checks.min_confidence.sections.<section> overrides it per section. Omitting the block, or setting low, admits everything, so the default configuration is behavior-preserving. checks.confidence_demotion additionally reports a low-confidence failing finding as a warning; it never promotes. Filtering happens in FinalizeSectionWithDiagnostics, before waiver auditing, suppression matching, and status computation, and therefore after the per-file findings cache: changing a threshold re-renders a scan rather than re-running one. SectionConfigHashes strips both settings through findingRelevantChecks so they stay out of every fingerprint family, including the conservative all-checks fallback. Removed findings are never silent. Each is counted per rule as confidence_filtered, per section as confidence_filtered_count, and listed with reason "confidence" under --include-suppressed. Confidence is deliberately not folded into Suppressed() or suppression_ratio, which keep meaning "findings teams work around". Demotion leaves all three fingerprints untouched, so a demoted finding still matches its baseline entry. Section keys are the ids that appear in scan output and in checks.disabled. Those diverge from the runner's registry ids (supply_chain vs supply-chain), so core.SectionKeys() holds the report-facing list and a test asserts every section a real scan reports is a configurable key. Text output sorts by confidence within each rule group with a stable sort; section finding order is untouched, so JSON and SARIF keep exact scan order. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A config checked into a repository gave no indication of which codeguard
produced it. WriteFile now records the writing release as a top-level
codeguard_version, so codeguard init and every SDK write path stamp it.
Stamping happens at the write boundary rather than in ApplyDefaults, so a
loaded config reports what its file recorded instead of the running binary;
that is what makes the field useful for spotting a config produced by a
different release. Rewriting refreshes it.
The stamp is provenance, not a compatibility gate. It is never validated, so
an absent, older, newer, or unparseable value always loads, and decoding was
already non-strict, so older binaries ignore the field. It is also absent from
SectionConfigHashes, so a release upgrade cannot discard cached findings; a
test pins that. It does reach ConfigHash, which only feeds waiver-audit history
dedupe, where recording a new observation after a config change is correct.
Release builds write v{{.Version}} via the GoReleaser ldflag, so the repository
config and the shipped JSON example use the v-prefixed form.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CI lint caught four staticcheck findings in the new test files: two WriteString(fmt.Sprintf(...)) calls that should be fmt.Fprintf (QF1012), and two negated conjunctions that read better under De Morgan (QF1001). Also drops a hand-rolled integer formatter in favor of strconv.Itoa. These were missed locally because golangci-lint suppresses staticcheck findings in repository code when its typecheck pass fails, and it fails here on the Homebrew Go 1.27.1 stdlib since the binary is built with go1.26.3. The local run therefore exited 0 while CI reported four issues. Pinning GOTOOLCHAIN to the linter's build version reproduces CI exactly and now reports 0 issues; the knowledge note records the corrected command. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
alxxjohn
added a commit
that referenced
this pull request
Sep 4, 2026
🤖 I have created a release *beep* *boop* --- ## [1.10.0](v1.9.0...v1.10.0) (2026-09-04) ### Features * **config:** add a minimum-confidence policy for findings ([53c7e5d](53c7e5d)) * **config:** minimum-confidence policy and config version stamp ([#178](#178)) ([d9def6e](d9def6e)) * **config:** stamp the codeguard version into written configs ([7d8ccd4](7d8ccd4)) --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Two config-level features plus the design work queued behind them.
Minimum-confidence policy turns the confidence every finding already carries into something a team can act on. Confidence has existed since findings gained the field —
high,medium,low, set by 177 call sites, normalized inNewFinding, emitted as a SARIF property — but nothing consumed it. Structural analyses (source-to-sink taint, the tree-sitter rule paths) sethighwhile regex line scans leave it unspecified, so a team wanting only the trustworthy half had to disable rules outright or baseline the noise. Nowchecks.min_confidencefilters,checks.confidence_demotionsoftens, and both are fully accounted for.Config version stamp records which codeguard wrote a config file, so a config checked into a repository carries its own provenance.
Planning docs for the next three accuracy/performance features: tree-sitter scan scale, cross-file taint, and this confidence policy.
Defaults are unchanged in every respect — an existing config produces byte-identical findings, counts, statuses, and exit codes.
Minimum-confidence policy
min_confidence.defaultdrops findings below a level;min_confidence.sections.<section>overrides it for one section. Omitting the block, or settinglow, admits everything.confidence_demotion: truereports alow-confidence finding on a failing rule as a warning instead. It never promotes and never touchesmediumorhigh. Off by default.Where it runs, and why that matters
Filtering happens in
FinalizeSectionWithDiagnostics— after diff scoping, before waiver auditing, suppression matching, and section status. That is after the per-file findings cache, which gives the property that makes thresholds cheap to tune: changing a threshold re-renders a scan rather than re-running one.Holding that property took one non-obvious fix.
SectionConfigHashesbuilds a fingerprint per config family, and its conservative""all-checks fallback hashescfg.Checkswholesale — so a new field underCheckswould have invalidated cached findings for every section outside the named families on every threshold change. Both settings are now stripped throughfindingRelevantChecksbefore fingerprinting, and a test asserts every family hash is unchanged while a genuinely finding-relevant setting (parsers.treesitter) still moves it.Nothing disappears quietly
A filtered finding is counted three ways: per rule as
confidence_filteredin therule_statsartifact, per section asconfidence_filtered_count, and listed with reasonconfidenceunder--include-suppressed. It is deliberately not folded intoSuppressed()orsuppression_ratio— a suppression records a team accepting a finding, a confidence filter records the scanner not being sure enough to show it, and collapsing the two would corrupt a metric that means "findings teams work around". A test assertsemitted + confidence_filtered + suppressed == total.Identity is untouched
Demotion rewrites
Level/SeverityafterNewFinding, so exact, context, and content fingerprints are all unchanged and a demoted finding still matches its baseline entry. Tests pin this in both directions.Section keys
The policy keys on the ids that appear in scan output and in
checks.disabled. Those diverge from the runner's internal registry ids — supply chain finalizes assupply_chainbut the registry calls itsupply-chain— so keying on the registry id would have made that section silently unconfigurable.core.SectionKeys()holds the report-facing list, andTestReportedSectionIDsAreConfigurableKeysasserts every section a real scan reports is a configurable key, so the divergence cannot reappear unnoticed.Ordering
Findings today have no explicit sort; they arrive in file-scan order. Sorting
section.Findingswould therefore have silently reordered JSON and SARIF. Instead the confidence sort is a stable sort within each text finding group: group identity and order are unchanged, equal-confidence findings keep scan order, and machine surfaces keep exact scan order. Pinned by a test that asserts JSON ordering is untouched.Config version stamp
config.WriteFile, socodeguard initand every SDK write path record the writing release. Verified end to end against the built binary.ApplyDefaults: a loaded config reports what its file recorded rather than the running binary, which is what makes the field useful for spotting a config produced by a different release. Rewriting refreshes it.SectionConfigHashes, so a release upgrade cannot discard cached findings (test included). It does reachConfigHash, which only feeds waiver-audit history dedupe, where a new observation after a config change is correct.v-prefixed form that release builds write via the GoReleaser ldflag.Verification
default: lowproducereflect.DeepEqualsections through a real scan.make codeguard-ciis 33 findings / 0 fail — exactly the count atmain, confirmed by diffing finding locations against a clean worktree at HEAD. (The repo's ownquality.ai.narrative-commentrule caught one of my doc comments during this; it is rewritten.)tests/support,tests/core,tests/report,tests/codeguard,internal/codeguard/runner/....gofmtandgo vetclean.tests/coreandtests/report; the latter drives rendering throughcodeguard.WriteReport.Pre-existing failure, not from this branch
TestSecuritySemanticAnalyzerScansNodeModulesWithinTargetfails onmainat 1205484 withSecurity status = "pass", want "fail", reproduced in a clean worktree at HEAD. ItsrequireTypeScriptSemanticRuntimeguard does not skip, so this is a real failure rather than a missing-runtime skip. It plausibly relates to the vendored/node_modulesscan bounding in 24cb12c but has not been diagnosed. Worth triaging separately — it meansgo test ./...is red before this branch.Local lint caveat
golangci-lint runlocally reports 2typecheckissues insidecrypto/internal/randutil/math/rand/v2because its bundled Go is older than the Homebrew toolchain (1.27.1). Both are stdlib paths; no repo code is implicated. CI's pinned version is the real gate.Docs
docs/checks.mdgains a confidence-policy reference,docs/features.mdgains YAML and JSON config examples for both features, andREADME.mdmentions the policy. Knowledge captured in.claude/knowledge/covers three things that would otherwise be re-learned: finding identity excludes presentation, post-cache policy must be stripped from cache fingerprints, and the section-id divergence.Follow-ups, deliberately not in scope
codeguard doctornote when the stamp differs from the running binary is the natural next step.rank; it stays a result property so external consumers see no schema churn.🤖 Generated with Claude Code